Back

BMC Genomics

Springer Science and Business Media LLC

Preprints posted in the last 90 days, ranked by how well they match BMC Genomics's content profile, based on 406 papers previously published here. The average preprint has a 0.29% match score for this journal, so anything above that is already an above-average fit.

1
The epigenomic landscape of deep lineage divergence: The case of the European sea bass

Longo, A.; Babbucci, M.; Jiao, Z.; Ferraresso, S.; Franch, R.; Bortoletti, M.; Bertotto, D.; Faggion, S.; Ilsley, G. R.; Papadogiannis, V.; Manousaki, T.; Kristoffersen, J.; Tsigenopoulos, C. S.; Macqueen, D. J.; Bargelloni, L.

2026-07-28 evolutionary biology 10.64898/2026.07.27.738391 medRxiv
Top 0.1%
38.1%
Show abstract

BackgroundUnderstanding the role of non-coding genomic variation in speciation remains a major challenge in evolutionary biology. Here, we investigated whether regulatory elements contribute to this process between Atlantic and Mediterranean lineages of European sea bass (Dicentrarchus labrax), a well-characterized case-study near speciation where barriers to introgression exist in the presence of connectivity between diverging populations. ResultsWe generated a novel, highly contiguous genome assembly, which was annotated at the epigenomic level using ATAC-seq and ChIP-seq with six embryonic developmental stages and five tissue types in adult fish, identifying thousands of promoters, enhancers, and open chromatin regions. Integrating this annotation with whole-genome sequence data from 65 individuals across three geographically distinct populations, we identified 57,505 outlier SNPs and 332 structural variants (SVs) showing elevated differentiation between Atlantic and East Mediterranean lineages. Outlier SVs affected key regulatory elements and coding genes, while outlier SNPs were enriched in regulatory elements, particularly enhancers active in adult tissues. Local genomic divergence correlated positively with regulatory element density, especially on chromosomes 1, 9, and 18, which are enriched in genes related to osmoregulation, immune response, and oxidative stress -- processes relevant to adaptation across contrasting marine environments. ConclusionsThese findings support a major role for regulatory variation in driving deep lineage divergence through local adaptation.

2
Comparative genomic and transcriptomic analyses of Strongylus vulgaris reveal developmental and evolutionary deployment of parasitism in a migratory equine nematode

Ripley, N. E.; Ripley, B.; Li, K.; Hudson, E. E.; Howe, D. K.; Kalbfleisch, T.; Smith, M.; Nielsen, M. K.

2026-07-22 genomics 10.64898/2026.07.17.739161 medRxiv
Top 0.1%
38.0%
Show abstract

BackgroundStrongylus vulgaris is a highly pathogenic equine strongyle whose larval stages migrate through the mesenteric arterial system, yet the molecular basis of its development, host association, and evolutionary biology remains poorly resolved. We generated an integrated genomic and transcriptomic resource to characterize genome structure, gene annotation, stage-associated expression, candidate secreted proteins, isoform diversity, gene-family evolution, putative horizontal gene transfer, and drug resistance-associated homologs in S. vulgaris. ResultsPacBio HiFi sequencing produced a 329.0 Mb genome assembly comprising 3,913 contigs, with a contig N50 of 147 kb and 94.8% BUSCO completeness, substantially improving the prior fragmented draft. Repeat annotation identified 127.5 Mb of repetitive sequence, representing 38.72% of the genome and dominated by unclassified repeats. Integration of RNA-seq-guided annotation with PacBio Iso-Seq evidence refined 16,387 loci and 24,546 transcripts, generating an isoform-retaining discovery proteome of 23,363 predicted proteins. Functional annotation supported 93.5% of predicted proteins and identified 1,530 unknown or weakly annotated candidates. Consensus secretome prediction identified 2,210 high-confidence putative secreted proteins. Stage-associated transcriptomics showed that development was the dominant axis of expression variation, with 5,299 genes differentially expressed between larvae and adults and strong adult sex-associated divergence. Isoform analysis identified 490 high-confidence isoform switches, concentrated primarily in the ML5 female-to-adult female transition. Comparative genomics identified 27,641 orthogroups, 201 S. vulgaris-specific orthogroups, and, after repeat-aware filtering, 48 expanded and 250 contracted gene families. Structure-guided annotation prioritized migratory-stage-enriched secreted candidates, including Cysteine-rich secretory proteins, Antigen 5, and Pathogenesis-related 1 (CAP), Sperm-coating protein (SCP) Tpx-1/Ag5/PR-1/Sc7 protein superfamily (TAPS) -like, von Willebrand factor type A (VWA) - domain, lipid-binding-like, DNase II-like, and peptidase-like proteins. Conservative screening retained 17 putative horizontal gene transfer (HGT) candidates, and drug resistance-homolog analysis recovered 18 nonredundant S. vulgaris homologs without canonical {beta}-tubulin benzimidazole-resistance substitutions. ConclusionsThese results establish the first integrated, high-quality molecular framework for S. vulgaris and show that its parasitic biology is developmentally structured, isoform-rich, and shaped by both conserved strongylid features and lineage-specific gene-family change. This resource provides a foundation for future studies of interhost migration, host interaction, parasite evolution, and genomic surveillance.

3
Genomic Inbreeding and Selection Signatures analyses in the Doberman Pinscher breed

Mulim-McCarthy, H.; Fragomeni, B.; Liu, S.; Rojas de Oliveira, H.

2026-06-08 genomics 10.64898/2026.06.04.730131 medRxiv
Top 0.1%
28.9%
Show abstract

The Doberman Pinscher population has undergone strong artificial selection for morphology and behavior, which can reduce genomic diversity and increase autozygosity. Here, we characterized the genome structure and identified selection signatures in Doberman Pinschers using complementary within- and between-population approaches. Genotypes from 3,226 Dobermans Dogs (Illumina CanineHD; 216,184 SNPs) provided by the Doberman Diversity Project were analyzed after purpose-specific quality control. Genomic inbreeding was quantified using four allele-frequency-based metrics and the runs of homozygosity (FROH) approach. Selection signatures were detected using intrapopulation (i.e., Runs of Homozygosity--ROH; Integrated Haplotype Score--iHS; and Number of Segregating Sites by Length--nSL) and interpopulation methods (i.e., Fixation Index--FST; Cross-Population Extended Haplotype Homozygosity--XP-EHH; and Cross-Population Number of Segregating Sites by Length--XP-nSL) comparing the Doberman Pinscher breed to Labrador Retriever (n=237). Dobermans showed high overall inbreeding, with a mean FROH of 0.42 (range 0.22-0.68), whereas the allele-frequency-based inbreeding estimators had similar means ([~]0.04). The partitioning of the ROH indicated high contributions from medium-to-long ROHs, consistent with recent inbreeding. The ROH scans identified 39,512 SNPs in ROH islands ([≥]50% frequency across individuals), with notable concentrations on CFA2, CFA3, and CFA31. Haplotype-based scans identified 2,820 candidate iHS SNPs and 2,173 candidate nSL SNPs (|score|>2). A common set of 310 SNPs was shared among ROH, iHS, and nSL, mapping near 279 genes that were mostly enriched for developmental pathways, particularly neurodevelopment and neuron-related cellular components. Between breeds, 349 highly differentiated SNPs were detected by FST, while XP-EHH and XP-nSL highlighted over 1,000 of Doberman-specific haplotype signals. A total of seven SNPs overlapped across FST, XP-EHH, and XP-nSL, which were located mainly on CFA8 ([~]59.48-60.61 Mb) near the KCNK10, SPATA7, PTPN21, NEGR1, and BTG1 genes. These genes are mainly linked to neural development and signaling, but BTG1 has also been associated with cardiomyocyte cell-cycle regulation, and KCNK10 with cardiac excitability and remodeling. Overall, the Doberman Pinscher breed exhibits high genome-wide autozygosity and levels of inbreeding. In addition, our results showed consistent, multi-method evidence of selection at loci associated with neurodevelopmental and regulatory pathways. These findings provide prioritized targets for follow-up studies that integrate phenotypes relevant to breed health and performance.

4
Variant intolerance scores in cattle

Lanigan, S.; Derks, M. F.; Johansson, A. M.; Johnsson, M.

2026-07-30 genetics 10.64898/2026.07.28.741161 medRxiv
Top 0.1%
26.5%
Show abstract

Variant intolerance methods score the essentiality of genes based on large datasets of genetic variants and have been used in population genomics of humans and model organisms. In this paper, we estimated Residual Variation Intolerance Scores for protein-coding genes and predicted protein domains in cattle. In agreement with results from other species, the most variant-tolerant genes and domains included genes related to olfaction and adaptive immunity, whereas the least-variant tolerant genes and domains included genes involved in fundamental cellular processes. There was a moderate positive correlation with estimates from orthologous human genes. We provide estimates of variant intolerance for cattle may be useful for genomic analyses of deleterious variants and population genomics in cattle.

5
Genome-wide meQTL mapping in cattle blood reveals cis and trans regulation of DNA methylation

Fouere, C.; Costes, V.; Besnard, F.; Le Danvic, C.; Patry, C.; Fritz, S.; Boussaha, M.; Jouin, M.; Boichard, D.; Kiefer, H.; Costa Monteiro Moreira, G.; Sanchez, M.-P.

2026-07-08 genetics 10.64898/2026.07.07.736355 medRxiv
Top 0.1%
26.2%
Show abstract

Background Complex traits are influenced by numerous variants, most of which have regulatory effects on gene expression that can be mediated by DNA methylation. Molecular QTL mapping is an approach that aims to dissect these effects. However, obtaining molecular phenotypes on a large scale is challenging, particularly in livestock species. In cattle, an epigenotyping array called EpiChip has recently been developed in the European RUMIGEN project. The EpiChip, which contains 43,317 CpG sites distributed all over the bovine genome, enables large-scale measurement of DNA methylation. This study aims to characterize the genetic determinism of blood DNA methylation in cows by estimating heritability and mapping cis- and trans-methylation QTLs (meQTLs). Results Whole blood samples from 4,457 genotyped Holstein cows were epigenotyped. Across all CpG sites, the heritability estimates averaged 24.6%. The local meQTL mapping at sequence-level for variable CpG sites (SD > 2.5%; n = 28,806) detected cis-meQTLs for 80.1% of the CpG sites, with sentinel SNPs located close to their associated CpGs. A two-step analysis was also conducted to identify long-range associations, with a particular focus on trans-meQTL hotspots. First, we identified CpG-SNP trans-associations using medium-density genotypes (50k SNPs) that revealed 31,846 SNPs with significant effects on 1 to 530 trans-CpG sites. Then, regions associated with at least 34 independent trans-CpGs were retained defining 31 hotpots. For each hotspot, a local sequence-level GWAS was conducted using the first principal component derived from the associated trans-CpGs. Out of the 31 detected hotspots, three were located close to transcription factor genes (RUNX1, NFIC and FOXA3) for which the associated trans-CpGs were enriched for the corresponding binding motif. Two other hotspots were located within KDM5A and KDM5B, and their corresponding trans-CpGs were strongly overrepresented in H3K4me3 narrow peaks in blood as well as in other tissues. Conclusions By identifying functional candidate genes associated with blood DNA methylation in cattle, these findings provide new insights into the regulatory architecture of DNA methylation in mammals, highlighting the value of large-scale molecular data from livestock populations.

6
Neuronal Gene Architecture in Cancer borealis Revealed by Long-Read Genome Assembly and Deep Transcriptomic Analysis

Raju, M.; Northcutt, A. J.; Schulz, D. J.

2026-08-20 genomics 10.64898/2026.08.12.744261 medRxiv
Top 0.1%
25.8%
Show abstract

Understanding the underlying neuronal function in non-model organisms requires accurate resolution of gene structure and transcript diversity. Here, we present a comprehensive genome annotation tor the Jonah crab (Cancer borealis), a key experimental system in crustacean neurobiology, with a particular focus on transcriptome-supported neuronal gene architecture. By integrating long-read genome assembly with extensive transcriptomic evidence, we reconstructed gene models with high confidence, enabling detailed characterization of exon-intron organization, alternative splicing, and isotorm diversity across gene families. Functional classification revealed extensive representation of neural-associated gene classes, including ion channels and receptors, transporters, enzymes, zinc finger proteins, histones, structural proteins, and cell adhesion molecules, alongside a large set of previously uncharacterized genes. In this study we particularly focused on the neuronal and ion channel gene families known to underlie circuit-level neuronal function in C. borealis. We provide an in-depth analysis of 87 genes spanning 17 neural-related gene families and 41 neuropeptides, detailing chromosomal localization, gene length, exon-intron configuration, and transcript-supported isotorm structure. For many of these genes, transcriptomic data confirmed expression and refined coding boundaries. Comparisons with existing transcriptomic datasets demonstrate strong concordance in gene expression patterns while also revealing novel transcripts and expanded gene family members not previously annotated. Together, this genome and transcriptome-integrated annotation establishes a high-resolution framework tor studying neuronal gene organization in C. borealis. T his resource enables direct connections between gene architecture, transcript diversity, and neural function, supporting future investigations in crustacean neurogenomics, comparative genomics, and the evolution of nervous system complexity.

7
Systematic assessment of sequencing depth requirements for Hi-C-derived metrics

Granovsky, A.; Polovnikov, K. E.

2026-08-01 bioinformatics 10.64898/2026.07.28.741297 medRxiv
Top 0.1%
23.3%
Show abstract

BackgroundHi-C experiments produce genome-wide chromatin contact maps from which structural features can be quantified across multiple genomic scales, ranging from megabase-scale compartments to kilobase-scale boundaries and chromatin loops. Although high-resolution analyses commonly rely on hundreds of millions of sequenced read pairs, the minimum sequencing depth required for different classes of Hi-C-derived features has not been systematically established. We therefore sought to determine these depth requirements systematically. ResultsWe performed progressive random subsampling of twelve Hi-C libraries representing multiple cell types and experimental protocols. Loop-density and loop-size inference from the P(s) log-derivative was evaluated across all libraries, whereas compartment, insulation, and boundary analyses were performed on a subset of nine libraries with comparable full-depth coverage. Rather than defining sufficient depth as an arbitrary fraction of the full-depth value, we introduced a biologically motivated criterion based on the variability between independent biological replicates. The required sequencing depth for each metric was defined as the point at which subsampling scatter first reached the variability observed between independent biological replicates, representing the accuracy that additional sequencing cannot improve upon. Using this criterion, loop-density estimation reached its biological floor at approximately 10 million read pairs, loop-size estimation at approximately 20 million, insulation scores and boundary detection at approximately 30 million, and compartment eigenvectors at 100-kilobase resolution at approximately 60 million read pairs. ConclusionsDifferent Hi-C-derived metrics require substantially different sequencing depths to achieve biologically meaningful accuracy. For all metrics, depth-induced variability fell below biological replicate variability well before full sequencing depth was reached. These thresholds provide practical guidance for experimental design and sequencing budget allocation, suggesting that, in many studies, increasing the number of biological replicates is likely to improve reproducibility more effectively than sequencing individual libraries to greater depth.

8
An endogenous retrovirus insertion disrupting bovine ALKBH8 causes a failure-to-thrive syndrome with immunodeficiency associated with juvenile mortality in Brown Swiss cattle

Glatthard, S.; Kadri, N. K.; Seefried, F. R.; Voitl, L. R.; Weber, B. A.; Schwarzenbacher, H.; Meister, S. L.; Gurtner, C.; OGrady, J. F.; Osbahr, M.; Leonard, A. S.; Meylan, M.; Pausch, H.; Droegemueller, C.; Jacinto, J.

2026-07-10 genomics 10.64898/2026.07.09.737535 medRxiv
Top 0.1%
23.1%
Show abstract

The Brown Swiss (BS) cattle breed is one of the major Swiss dairy breeds. Intensive selection and the widespread use of few elite sires in artificial insemination have increased inbreeding and the occurrence of deleterious recessive alleles in the homozygous state. Analyzing life trajectories in large, genotyped cohorts can identify hidden recessive disorders that are difficult to detect using traditional case-control association testing. Long-read DNA sequencing enables precise detection of causal alleles, including structural variants. This study aimed to (1) identify cryptic recessive loci affecting rearing performance in Swiss BS cattle, (2) evaluate their impact on survival, (3) characterize the associated phenotype, (4) identify the causal variant using long-read whole-genome sequencing, and (5) assess its functional impact. Using Homozygous Haplotype Enrichment/Depletion (HHED) mapping, we identified a risk haplotype (BH39) on chromosome 15 spanning from 16,276,819 bp to 16,446,984 bp that was associated with increased juvenile mortality within the first 180 days of life when present in the homozygous state. The BH39 occurred at a frequency of approximately 4.5% in Swiss BS cattle and 5.3% in German and Austrian BS cattle, and homozygous carriers exhibited a significantly reduced first-year survival rate. Five females homozygous for BH39 underwent clinical examination. They all showed recurrent respiratory disease, impaired growth, poor body condition, rough hair coat, and brown-discolored teeth. Pathological examination revealed bronchopneumonia and eosinophilic enteritis. Clinicopathological findings indicated failure to thrive and immunodeficiency. Long-read WGS of two BH39 homozygous calves revealed a private homozygous coding variant that was in high linkage disequilibrium with BH39. The identified structural variant was an insertion of a large transposable element (10.4 kb ERVK[2-1-LTR]) into the third exon of ALKBH8 (NM_001080341.2 c.267_268indel). Full-length RNA sequencing of cerebellum and liver from a homozygous calf revealed that the endogenous retrovirus (ERV) insertion introduces a cryptic transcription termination signal, truncating ALKBH8 mRNA. This study demonstrates that exploring population-scale genomic data and mining thousands of life-history records, followed by veterinary follow-up evaluations and molecular genetic analyses, provides an effective strategy for identifying cryptic recessive disorders that shorten the lifespan of cattle. The findings provide strong evidence that the ERV insertion into the coding sequence of ALKBH8 represents a loss-of-function variant that causes a previously undescribed recessive disorder that results in increased rearing loss. Interpretive summaryWe identified a recessive disorder in Brown Swiss cattle that causes retarded growth, recurrent infections, immunodeficiency, and increased mortality during the first year of life. Using population-scale genomic data, clinical investigations, and long-read sequencing, we linked the disorder to an exonic transposable element insertion disrupting ALKBH8. The identification of the causal variant now enables direct genetic testing and the implementation of genome-based mating strategies to avoid carrier-by-carrier matings and, consequently, prevent the birth of affected homozygous offspring. We demonstrate the utility of integrating large-scale breeding records, veterinary phenotyping, and advanced genomics to identify hidden defects affecting livestock health and productivity.

9
Elevated isoform richness in males largely reflects transcriptional noise rather than proteomic complexity

Sherin, L. M.; Johnson, B. D.; Corral-Lopez, A.; van der Bijl, W.; Mank, J. E.

2026-08-18 genomics 10.64898/2026.08.10.744030 medRxiv
Top 0.1%
22.9%
Show abstract

Alternative splicing (AS) can generate multiple RNA isoforms from a single gene and is thought to contribute to phenotypic divergence, including differences between the sexes. Studies in several organisms have documented sex differences in splicing, however these have been largely reliant on short-read RNA sequencing which requires complex algorithms to assemble full-length transcripts and may underestimate both isoform diversity and sex differences in splicing. We used long-read, single molecule RNA-Seq to build a more complete catalog of sex-biased splicing in Poecilia reticulata, a focal species for studies of sexual dimorphism. Pairing long-read sequencing with isoform-level analyses, we identified a sixfold higher proportion of sex-biased splicing genes (37%) compared with short and long-read event-level approaches (6%, 11%). AS was common (70% genes) but only 54% of isoforms produced unique open-reading frames (ORFs). We found that males exhibited greater isoform richness than females in both tail and gonad tissues but produced a smaller proportion of isoforms with unique ORFs, suggesting that much of the increased isoform variation is unlikely to expand proteomic complexity and may instead reflect stochasticity during splicing rather than intentional transcriptional products intended for translation. Despite widespread AS, we found only 2.3% of genes exhibited sex-biased isoform switching, and only 52% of these switches generated distinct sex-biased ORFs. Together, our long-read data suggest that although isoform diversity is more extensive than previously appreciated, most alternative isoforms are unlikely to generate novel proteins. Instead, a relatively small number of sex-biased isoforms may disproportionately contribute to proteomic divergence between the sexes.

10
Genome-Wide Selection Signatures in Nili-Ravi Buffalo (Bubalus bubalis) Reveal a T-Cell Costimulatory and Cytokine-Signaling Gene Network Distinct from Classical Bovine Tuberculosis Candidate Genes

Ahmad, A.; bakar, A.; Laeeque, S. M.; Khan, W. A.; Kaul, H.; Manan, A.; mustafa, h.

2026-08-11 genomics 10.64898/2026.08.10.743898 medRxiv
Top 0.1%
22.4%
Show abstract

Genomic signatures of selection can reveal loci underlying adaptation and disease resistance in livestock populations, but such analyses in water buffalo (Bubalus bubalis) have historically been constrained by the absence of a chromosome-level, species-native reference genome for SNP array data. We re-analyzed genotype data from 85 Nili-Ravi buffalo (Axiom Buffalo Genotyping 90K array, originally positioned using bovine (Bos taurus, UMD3.1) proxy coordinates, by performing a full coordinate liftover to the buffalo-native UOA_WB_1 assembly using an independently published SNP remapping resource. Following quality control (51,209 markers retained), haplotype phasing, and genome-wide integrated haplotype score (iHS) and Wrights Fst (case/control) selection scans, we evaluated 14 classical bovine-tuberculosis (bTB) candidate genes and identified six additional genes with putative immune function through an unbiased genome-wide screen. None of the 14 classical candidates (including SLC11A1, the Toll-like receptors, and IFNG) reached genome-wide significance in either scan. In contrast, six novel loci TNFSF18, IL2RB, TNFRSF19, IRF2, IL15, and CD28 showed significant iHS or Fst signals, four of which (TNFSF18, IL2RB, IL15, CD28) converge functionally on T-cell costimulation and cytokine receptor signaling (KEGG pathways map04660 and map04060, Bos taurus proxy annotation). Using extended haplotype homozygosity (EHH) decay, haplotype furcation structure, and per-marker haplotype counts as three independent lines of corroborating evidence, we classified these six genes into confidence tiers: TNFSF18 and IL2RB showed the strongest, most balanced support, while CD28 and IL15 signals were driven by very few haplotypes (3 and 5 of 30, respectively) and should be interpreted cautiously pending replication. These findings suggest that adaptive, cell-mediated immune signaling rather than the innate/macrophage-centred mechanisms emphasized by existing bTB candidate gene panels may be a more productive avenue for future selection studies in Nili-Ravi buffalo, while underscoring the value of buffalo-native coordinate systems for accurate genomic inference in this species.

11
Pangenome Graph Node-Phenotype Association shows GWAS-like quality results with only few individuals

Carrette, C.; Sabot, F.; Muller, C.

2026-08-01 bioinformatics 10.64898/2026.07.31.741971 medRxiv
Top 0.1%
18.7%
Show abstract

PurposeWe introduce GO_SCPLOWRAC_SCPLOWNPA, standing for Graph Node-Phenotype Association, a method performing a GWAS-like analysis on a pangenome variation graph (PVG) built using a small number of individual genome sequences, without the need for additional population materials or kinship information for qualitative phenotypes. This method reduces the number of individuals required for association studies and prevents reference bias from variant calling in these types of analyses. BackgroundA PVG represents the multiple alignment of a set of complete genomes. It contains all variations, from single nucleotide polymorphisms (SNPs) to large structural variations (SVs), which are represented as nodes in the graph. By integrating phenotype information within nodes, we can assign a Phenotype Score (PS) to each node in the PVG and identify phenotype-related regions directly within it. These regions represent statistically significant shifts in PS distribution, highlighting their implication in the phenotype. Finally, GO_SCPLOWRAC_SCPLOWNPA provides their positions and scores for further analysis. ResultsThis method was tested using simulated data and two publicly available datasets: the Sub1A gene locus for Oryza sativa in a 13 individuals PVG, and the insertion responsible for the white-headed cattle with a PVG of 24 individuals. Source code of GO_SCPLOWRAC_SCPLOWNPA is available here https://forge.ird.fr/diade/graphgwas/granpa under GNU GPLv3. ConclusionGO_SCPLOWRAC_SCPLOWNPA was able to identify the expected area in two simulated datasets and the responsible loci for these two known traits using only a few dozen complete genomes in these PVGs. While currently limited to qualitative phenotypes, this method opens the way to more efficient ones relying on PVGs and few individuals.

12
Correction of the cytosine deamination artifacts in FFPE-based sequencing experiments

Płonka, W.; Kostka, D.; Lalik, A.; Kurpas, M.; Dinh, K. N.; Sitkiewicz, M.; Kimmel, M.; Rzyman, W.; Jaksik, R.

2026-08-19 bioinformatics 10.64898/2026.08.11.744151 medRxiv
Top 0.1%
18.7%
Show abstract

Formalin-fixed, paraffin-embedded (FFPE) tissues remain an essential resource for molecular studies, yet formalin-induced cytosine deamination introduces characteristic C>T/G>A artifacts that compromise the accuracy of next-generation sequencing (NGS) analyses. Numerous computational methods and enzymatic DNA repair strategies have been proposed to reduce these artifacts, but no systematic comparison across tools and experimental conditions exists. Here, we evaluate the performance of seven computational approaches (SOBDetector, Ideafix, MicroSEC, FFPolish, DeepOmics FFPE/FFPE-PLUS, FFPErase) together with the NEBNext(R) FFPE DNA Repair Mix v2, a multi-enzyme repair system applied during DNA preparation. Using three independent datasets, one based on whole genome sequencing (CGCI-BL) and two on whole exome sequencing (TCGA-PC and SUT-LUAD, the latter containing enzymatically repaired samples), and matched fresh-frozen samples as the gold standard, we assess precision, sensitivity, and artifact reduction efficiency across all methods. We further examine the potential synergy between enzymatic repair and post-sequencing computational filtering. Our results provide practical guidelines for FFPE artifact correction and demonstrate that enzymatic treatment provides the best results, while among the computational methods, FFPErase offers the most robust reduction of cytosine deamination artifacts while maximizing the retention of true somatic variants. KEY MESSAGESO_LIFormalin fixation in FFPE samples introduces artifacts that can significantly affect the accuracy of NGS analyses. C_LIO_LIAmong the evaluated approaches, enzymatic repair using NEBNext(R) FFPE DNA Repair Mix v2 achieves the most effective reduction of sequencing artifacts. C_LIO_LIComputational methods vary in performance, with FFPErase showing the most robust balance between artifact removal and retention of true somatic variants. C_LIO_LICombining enzymatic repair with computational filtering did not lead to consistent improvements in performance across datasets. C_LI

13
Long-read transcriptomics highlights venom gland specialization and Inhibitor Cystine Knot (ICK) rich toxin diversity in Philippine tarantulas

Ragasa, L. R. P.; Dumbrique, M. M. U.; Gamboa, S. A. S.; Baile, A. G. M.; Acuna, D. C.; Frisco-Cabanos, H. L.; del Rosario, R. C. H.; Guevarra, L. A.; Santiago-Bautista, M. R.

2026-07-18 bioinformatics 10.64898/2026.07.14.737467 medRxiv
Top 0.1%
18.6%
Show abstract

Animal venoms are a rich source of bioactive molecules, yet their diversity remains incompletely characterized in many species. Here we present the first long-read transcriptomic analysis of venom glands from Philippine tarantulas (Theraphosidae), a highly endemic but understudied group. Using Oxford Nanopore sequencing, we reconstructed near full-length venom gland transcriptomes across multiple species and identified extensive repertoires of toxin-encoding peptides. Venom glands were enriched in cysteine-rich inhibitor cystine knot (ICK) peptides, which dominated the toxin landscape and are known modulators of ion channels. Cross-species comparative analyses revealed a distinct transcriptional signature separating venom from non-venom tissues, driven by coordinated expression of toxin-associated and regulatory gene families. Phylogenomic reconstruction based on orthologous peptides recovered expected taxonomic relationships while revealing potential lineage-specific diversification and potential cryptic taxa. Despite a conserved core set of toxin families, substantial variation in toxin composition was observed among species, consistent with rapid evolution driven by gene duplication and functional divergence. Analysis of highly expressed ICK peptides showed a conserved cysteine framework alongside marked sequence variability in inter-cysteine regions, supporting a model in which structural stability is maintained while functional diversification proceeds. Together, these findings establish the first long-read transcriptomic resource for Philippine theraphosid spiders, reveal a conserved molecular signature underlying venom gland specialization, and provide new insights into the diversification of ICK toxin repertoires that may facilitate future evolutionary and functional studies, including the discovery and characterization of bioactive venom peptides.

14
A gapless telomere-to-telomere reference genome of Ostreococcus tauri RCC4221 with expanded annotation of medium-sized ncRNAs

Liu, G.; Bousquet, L.; Mayeur, H.; Manirakiza, E.; Daric, V.; Klopp, C.; Noirot, C.; Lopez-Escardo, D.; Grimsley, N. H.; Yau, S.; Krasovec, M.; Echeverria, M.; PIGANEAU, G.

2026-07-14 genomics 10.64898/2026.07.10.737489 medRxiv
Top 0.1%
18.5%
Show abstract

Marine photosynthetic microbes contribute substantially to global primary production, yet many algal lineages still lack reference genomes with the continuity and annotation quality required for fine-scale structural, regulatory and comparative analyses. Ostreococcus tauri, one of the smallest known free-living photosynthetic eukaryotes, has been a model marine picoeukaryote for over two decades. Despite successive improvements to its historical reference genome, previous assemblies retained hundreds of gaps and incomplete genes, hampering high-resolution genomic analyses. Here, we present O. tauri RCC4221 genome version 2026, a telomere-to-telomere assembly of all 20 chromosomes spanning 13.34 Mb with no gaps. This assembly combines PacBio long-read sequencing, Illumina short-read polishing, correction of unresolved regions guided by independent Nanopore-based assemblies. The updated reference supports a curated annotation comprising 7,683 protein-coding genes, 48 tRNA genes, 3 rRNA operons, 116 medium-sized noncoding RNAs, one signal recognition particle RNA and 138 small nucleolar RNAs. It also improves gene-model integrity and recovers candidate coding loci absent from the 2014 reference. Structural analyses resolved the organization of the two atypical low-GC chromosome 2 and 19 that contain duplicated regions that were collapsed or misrepresented in previous assemblies. Finally, bisulfite sequencing and PacBio SMRT sequencing revealed a dual DNA methylation landscape, with CG-context cytosine methylation concentrated in gene bodies and N6-methyladenosine (m6A) enriched at the start codon. The updated O. tauri 2026 assembly provides a complete and curated reference resource for chromosome biology, comparative genomics, epigenomics and RNA biology in a model marine picoeukaryote.

15
The Blood RNA Stability Atlas: defining temporal structure and trait-state programs in the human whole-blood transcriptome

Baltazar, W. C.; Messing, R. O.; Ferguson, L.

2026-07-08 genomics 10.64898/2026.07.07.737044 medRxiv
Top 0.1%
18.3%
Show abstract

Whole-blood RNA biomarkers are widely used for diagnosis and disease monitoring, but their utility depends not only on abundance but also on temporal stability, a property that is not routinely incorporated into biomarker design. We analyzed 968 longitudinal whole-blood transcriptomes from 165 healthy individuals across eight independent studies spanning diverse platforms, time scales (50 minutes to 16 weeks), and common environmental exposures. Using a cross-study analytical framework integrating variance partitioning, repeatability, and time-associated differential expression, we quantified temporal stability for 6,064 RNAs and classified transcripts into "trait" (stable) and "state" (dynamic) categories representing the extremes of longitudinal changes in transcript abundance. We identified 1,118 trait RNAs exhibiting stable within-individual levels of abundance but substantial inter-individual variability, enriched for whole-blood eGenes (P = 6.0 x 10-20), supporting a genetic basis for stability. In contrast, 1,504 state RNAs showed context-dependent temporal variation and were enriched for translation and RNA-binding pathways. Integration with genetic datasets revealed that 4,395 (72%) blood transcripts were linked to at least one whole-blood eQTL, collectively associated with 18,358 GWAS trait relationships, providing disease-relevant context for transcript stability. We developed the Blood RNA Stability Atlas to integrate these features and demonstrate both top-down (disease-to-gene) and bottom-up (gene-to-context) applications for biomarker prioritization and interpretation. These findings establish temporal stability as a defining property of the blood transcriptome and provide a practical, publicly accessible framework for distinguishing stable baseline abundance levels from context-dependent transcriptional responses, informing biomarker selection, study design, and hypothesis generation.

16
A Cesium Chloride Gradient Ultracentrifugation-Based Method for the Isolation of DNA from Diverse Recalcitrant Plant Species for Nanopore Sequencing

Labbancz, J.; Dhingra, A.

2026-08-21 molecular biology 10.64898/2026.08.18.745475 medRxiv
Top 0.1%
18.3%
Show abstract

Developments in Nanopore sequencing have enabled telomere to telomere genomic assembly as a routine technique in genomic research. Nanopore DNA sequencing for genomic assembly is typically performed on native DNA molecules, making it particularly sensitive to the quality of input DNA, with contaminating molecules limiting data yields and reducing read quality. As pangenome analysis gains interest, particularly in non-model plant species which are often rich in inhibitory secondary metabolites, the development of methods which can improve the quality and throughput of nanopore sequencing is essential. Here we describe a method for isolation of total DNA from the leaf tissues of diverse Viridiplantae species. The initial lysis buffer consists of a modified CTAB buffer, incorporating dimethyl sulfoxide for the reduction of viscosity, which can be problematic in many plant DNA preparations. An organic extraction with 2-butoxyethanol is utilized to further extract phenolic compounds which may be sufficiently hydrophilic to evade chloroform extraction, while reducing aqueous phase volume. Further cleanup via cesium chloride (CsCl) ultracentrifugation is performed to minimize the carryover of residual contaminating macromolecules. Samples prepared using this method are of consistent high quality, even when extracted from challenging late season leaf tissue or secondary metabolite rich species. Sequencing results from samples prepared by this method outperform those obtained from typical modified CTAB DNA isolation techniques in both quantity and quality. We tested sequencing performance from Vitis DNA isolated using a modified CTAB method and Vitis DNA isolated using the CsCl ultracentrifugation-based method described here. DNA isolated via the method described here produced 83% more >Q10 sequence data (52.61 Gb vs. 28.8 Gb), resulted in a 60% greater read N50 despite more handling steps (32.78kb vs. 20.45kb), and resulted in a higher modal read quality (Q27 vs. Q24). The consistency of this method across diverse plant taxa suggests its use as a general method for DNA isolation prior to Nanopore sequencing and genomic assembly for diverse plant taxa.

17
BLink-seq delivers population-scale haplotypes without long reads: a scalable framework for non-model genomics

Iqbal, A. R.; Dimens, P. V.; Rick, J. A.; Munn, P. R.; McNairn, A. J.; Landis, J. B.; Schembri, R.; Chan, Y. F.; Kucka, M.; Therkildsen, N. O.; Grenier, J. K.

2026-08-07 genomics 10.64898/2026.08.03.742036 medRxiv
Top 0.1%
18.3%
Show abstract

Information about segregating haplotypes and structural variation (SV) can be extremely rich for a variety of applications in population genomics but remains largely inaccessible for many non-model species. Of the available methods, linked-read sequencing is especially promising for its low cost and scalability, but its adoption remains limited. One existing linked-read method is Haplotagging, which barcodes sequencing reads to reconstruct long molecules that encode haplotype information, with the potential to generate phased whole-genome data and detect structural variants. In this study, we present BLink-seq, a novel Haplotagging method that is compatible with standard short-read next-generation sequencing platforms, is locally reproducible with low-cost reagents, and is scalable for high-throughput sample processing. We optimized library preparation parameters, explored their relationship to linked-read library metrics, and validated phasing performance and structural variant detection in two evolutionary extremes: an experimental Drosophila melanogaster cross of inbred lines carrying known inversions, and four Atlantic silverside (Menidia menidia) parent-offspring trios sourced from highly outbred, wild-caught populations. We then applied our protocol to a cohort of 376 silversides to demonstrate its scalability and potential for SV detection and genotype imputation. Using BLink-seq, we generated chromosome-scale phased blocks and identified known inversions in both validation datasets. We discovered previously uncharacterized structural complexity within a known adaptive inversion on silverside chromosome 11, demonstrating that linked-read data can refine our understanding of SV architecture beyond what short reads alone can resolve. Finally, we provide a user guide for researchers interested in using BLink-seq.

18
Genomic insights into the karyotypic radiation of a narrow endemic holocentric plant Carex helodes

Gomez-Ramos, I.; Sanchez-Villegas, R.; Mohan, A. V.; Cornet, C.; Marques, A.; Maguilla, E.; Martin-Bravo, S.; Lucek, K.; Escudero, M.

2026-07-18 genomics 10.64898/2026.07.14.738159 medRxiv
Top 0.1%
18.1%
Show abstract

Holocentric chromosomes allow rapid genome changes through chromosomal rearrangements such as fissions, fusions, inversions or translocations. The plant genus Carex shows one of the highest rates of karyotypic evolution among holocentric organisms. We studied the genomic patterns underlying chromosomal rearrangements in the karyotypic radiation of the narrow endemic species Carex helodes (2n = 68-75). Comparing genome assemblies of C. helodes from the two karyologically distinct extremes of its European distribution, revealed a striking number of eight chromosomal rearrangements including fusions, translocations and inversions. Genomic breakpoints are gene-poor and TE-rich, corroborating findings in other species and suggesting common genomic characteristics that facilitate the evolution and establishment of chromosomal rearrangements. We identified a chromosomal inversion exhibiting patterns of purifying selection and enrichment in functional genes that potentially mediate rearrangement tolerance. Conversely, another inversion displayed elevated sequence divergence and enrichment in response to temperature stress and phosphate limitation, matching key environmental variables that differ between the study localities. The establishment of chromosomal rearrangements along Carex helodes European populations was likely driven by demographic bottlenecks and distinct genomic features at breakpoints. Our findings provide preliminary evidence on the rearrangement role in population differentiation either as reproductive barriers or as genomic islands of differentiation.

19
An evaluation of clustering and assembly strategies from Iso-Seq data in the absence of reference genomes in non-model animals

Eleftheriadi, K.; Vazquez-Valls, M.; Fernandez, R.

2026-07-08 evolutionary biology 10.1101/2025.09.18.677004 medRxiv
Top 0.1%
18.1%
Show abstract

Transcriptome assembly enables the recovery of expressed genes and isoforms, but the optimal strategy for reconstructing transcriptomes from long-read sequencing remains unresolved. In particular, establishing best practices for generating accurate gene models and selecting representative isoforms is essential for comparative genomics, as for orthology inference typically only the longest isoform per gene model is included. Here, we systematically compare clustering and de novo assembly methods using PacBio Iso-Seq data from 17 animal lineages spanning seven phyla, most of them non-model species, with the goal of investigating which methodology is more adequate to select one isoform per gene model, in the absence of specific pipelines to do so. We evaluate four approaches: isoseq cluster, CD-HIT, RNA-Bloom2 and isONform. We benchmark them with short-reads using Trinity, assessing assembly quality with BUSCO completeness, short-read mapping rates, coding sequence recovery, and longest isoform prediction. Our results show that CD-HIT clustering at high similarity thresholds ([≥]99%) yields the most complete and coding-rich long-read transcriptomes, rivaling Trinity while avoiding its high redundancy. Consensus-based methods such as isoseq cluster and isONform recover fewer single-copy orthologs (mirrored in a lower BUSCO score) and achieve lower mapping rates, while RNA-Bloom2 provide intermediate performance with reduced duplication. Together, these findings establish, to date, CD-HIT as a robust and practical strategy for transcriptome reconstruction from long-read data when genomic references are unavailable. By benchmarking de novo methods across a taxonomically broad dataset, this work defines the realistic capabilities of long-read transcriptome reconstruction in the absence of a reference genome and provides practical guidance for deriving high-quality gene models and selecting representative isoforms for orthology inference in non-model species.

20
The role of long-range transcriptional regulation in interpretation of non-coding variants associated with human disease

Mandic, K.; Hrsak, D.; Uljanic, F.; Lenhard, B.; Baresic, A.

2026-06-17 genomics 10.64898/2026.06.15.731051 medRxiv
Top 0.1%
17.8%
Show abstract

Genome-wide association studies (GWAS) are the key tools for the discovery of associations between single nucleotide polymorphisms (SNPs) and phenotypic traits and have been successfully applied to many diseases and disorders. However, a great challenge is to find the gene affected by the non-coding fraction of SNPs, especially if the gene is distal in terms of genomic distance. In this study, we present a novel approach, named targPred, which utilises genomic regulatory blocks (GRBs) for inference of a connection between a certain SNP/locus and the target gene located in the same GRB, in a more robust and generalisable manner. We identified that many disease traits such as cancer and psychiatric disease have a propensity for long-range regulation. Furthermore, we showcased a childhood obesity locus which is connected to the distal BDNF gene. Finally, we propose a new web-based service based on enhancer-promoter association, to facilitate finding the causal genes for a wide array of traits and conditions.